Papers with English data

58 papers
Token Sequence Labeling vs. Clause Classification for English Emotion Stimulus Detection (2020.starsem-1)

Copied to clipboard

Challenge: Emotion stimulus detection is the task of finding the cause of an emotion in a textual description.
Approach: They propose to evaluate whether clause classification or token sequence labeling is better for emotion stimulus detection in English.
Outcome: The proposed framework compares clause classification and token sequence labeling on four English datasets.
Exploring Cross-Lingual Voice Conversion Methods for Anonymizing Low-Resource Text-to-Speech (2026.eacl-short)

Copied to clipboard

Challenge: a growing number of speech synthesis systems clone a person's voice, a new study finds . a variety of voice conversion techniques can mask speaker identities in low-resource text-to-speech systems.
Approach: They compare voice conversion techniques to mask speaker identities in text-to-speech systems . they build and evaluate speaker-anonymized systems for two Canadian Indigenous languages .
Outcome: The proposed methods are compared with other approaches for using voice conversion to mask speaker identities in low-resource text-to-speech systems.
Multilingual Neural Semantic Parsing for Low-Resourced Languages (2021.starsem-1)

Copied to clipboard

Challenge: a large amount of training data is needed to understand multilingual semantic parsing models.
Approach: They propose to use machine translation to bootstrap multilingual training data from English data.
Outcome: The proposed model outperforms existing models on human-written sentences and the state-of-the-art models on the public NLMaps dataset.
Seeking Clozure: Robust Hypernym extraction from BERT with Anchored Prompts (2023.starsem-1)

Copied to clipboard

Challenge: Existing methods for extracting hypernym knowledge from large language models are unclear whether they fail due to a lack of knowledge or shortcomings.
Approach: They propose to use pattern-based hypernym extraction as a diagnostic tool to examine hypernomy knowledge encoded in BERT.
Outcome: The proposed method compares the results of two different methods on six English data sets and on challenge sets of rare and abstract concepts.
Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages (2025.naacl-short)

Copied to clipboard

Challenge: a large-scale text classification dataset encompassing 1504 languages is needed to address this gap . low-resource languages are often overlooked due to the scarcity of evaluation datasets.
Approach: They propose to use translations of the Bible to construct a large-scale text classification dataset that covers 1504 languages and annotate them using crowdsourcing.
Outcome: The proposed dataset covers 1504 languages and is available to the public.
A Survey on Multilingual Mental Disorders Detection from Social Media Data (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies on mental disorders focus on English data, overlooking critical signals that may be present in non-English texts.
Approach: They present a list of 108 social media datasets that can be used to train NLP models for mental health screening in 25 languages.
Outcome: The proposed datasets cover 25 languages and can be used to train models for mental health screening.
Don’t Use English Dev: On the Zero-Shot Cross-Lingual Evaluation of Contextual Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Multilingual contextual embeddings have demonstrated state-of-the-art performance in zero-shot cross-lingual transfer learning.
Approach: They show that English dev accuracy makes it difficult to obtain reproducible results . they recommend providing oracle scores alongside zero-shot results if possible .
Outcome: mBERT and XLM have shown strong performance on cross-lingual recognition, text classification, dependency parsing, and other tasks.
Cross-Lingual Transfer of Cognitive Processing Complexity (2023.findings-eacl)

Copied to clipboard

Challenge: Recent studies indicate that multilingual language models utilize structural similarities between languages to facilitate cross-lingual transfer.
Approach: They propose a multilingual model that uses structural similarities between languages to facilitate cross-lingual transfer by a meaningful bias towards sentence length and cross-linguistic differences.
Outcome: The proposed model can predict varied patterns for 13 languages, despite being fine-tuned only on English data.
Compositional Generalization in Multilingual Semantic Parsing over Wikidata (2022.tacl-1)

Copied to clipboard

Challenge: Semantic parsers are mostly designed for and evaluated on English resources, such as CFQ.
Approach: They propose a method for creating a multilingual, parallel question-query dataset . they analyze compositional generalization of parsers in Hebrew, Kannada, Chinese, and English .
Outcome: The proposed method analyzes compositional generalization of parsers in Hebrew, Kannada, Chinese, and English.
Negation Detection in Dutch Spoken Human-Computer Conversations (2022.lrec-1)

Copied to clipboard

Challenge: Existing negation detection methods in English are not available.
Approach: They propose to annotate a Dutch dialogue corpus with negation cues and their scopes.
Outcome: The proposed method can detect negation cues and scope in Dutch dialogues with high precision and recall.
Only for the Unseen Languages, Say the Llamas: On the Efficacy of Language Adapters for Cross-lingual Transfer in English-centric LLMs (2025.acl-srw)

Copied to clipboard

Challenge: Most state-of-the-art large language models (LLMs) are trained mainly on English data, limiting their effectiveness on non-English, especially low-resource, languages.
Approach: They train language adapters for 13 languages and evaluate their effectiveness on downstream tasks using either task adapters or in-context learning.
Outcome: The proposed language adapters improve performance for languages not seen during pretraining, but provide negligible benefit for seen languages.
Bhaasha, Bhāṣā, Zaban: A Survey for Low-Resourced Languages in South Asia – Current Stage and Challenges (2025.findings-emnlp)

Copied to clipboard

Challenge: a survey examines the current efforts and challenges of NLP models for South Asian languages . there are more than 650 languages in South Asia, but many have very limited computational resources or are missing from existing models.
Approach: a survey examines efforts and challenges of NLP for South Asian languages . they focus on transformer-based models such as BERT, T5, & GPT . findings highlight substantial issues, including missing data in critical domains .
Outcome: The findings highlight significant issues, including missing data in critical domains . the survey aims to raise awareness within the NLP community for more targeted data curation .
Parameter-Efficient Neural Reranking for Cross-Lingual and Multilingual Retrieval (2022.coling-1)

Copied to clipboard

Challenge: State-of-the-art neural rankers are notoriously data-hungry and rarely used in multilingual and cross-lingual retrieval settings.
Approach: They propose to use Sparse Fine-Tuning Masks and Adapters to transfer rankers trained on English data to other languages and cross-lingual setups by means of multilingual encoders.
Outcome: The proposed methods outperform standard zero-shot transfer with full MMT fine-tuning while being more modular and reducing training times.
DLIR: Spherical Adaptation for Cross-Lingual Knowledge Transfer of Sociological Concepts Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for identifying nuanced sociological concepts fail to capture domain-specific subtleties or require extensive parallel data.
Approach: a new approach to aligning nuanced sociological concepts is proposed . a dual-branch LoRA approach captures core semantics and counteracts specific language perturbations.
Outcome: a new approach outperforms baselines on cross-lingual sociological concept retrieval across 10 languages.
HISTOIRESMORALES: A French Dataset for Assessing Moral Alignment (2025.naacl-long)

Copied to clipboard

Challenge: HistoiresMorales is a dataset based on moralStories in French . it is based upon annotations of moral values within the dataset .
Approach: They propose a dataset in French that aims to align language models with moral values . they use annotations to ensure their alignment with French norms .
Outcome: The proposed dataset guarantees grammatical accuracy and adaptation to the French cultural context.
Multimodal Frame Identification with Multilingual Evaluation (N18-1)

Copied to clipboard

Challenge: FrameNet Semantic Role Labeling aims to disambiguate situations around predicates using textual representations.
Approach: They extend a frame identification task to leverage multimodal representations to improve FrameNet Semantic Role Labeling.
Outcome: The proposed system outperforms its unimodal counterpart on the English frameNet and its German counterpart on IMAGINED words.
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model (2025.acl-long)

Copied to clipboard

Challenge: Existing models for large vision-language tasks are trained on English data, which makes them struggle to understand non-English input and fail to generate output in the desired target language.
Approach: They conduct multi-stage experiments on 13 vision-language tasks and 43 languages . they find that one can include as many as 100 training languages simultaneously with as little as 25-50% of non-English data .
Outcome: The proposed model outperforms existing models in 14 tasks and 56 languages.
Reproducing and Regularizing the SCRN Model (C18-1)

Copied to clipboard

Challenge: Recurrent neural networks (RNNs) have demonstrated tremendous success in sequence modeling . naive dropout, variational dropout and weight tying are common techniques used to regularize the SCRN model .
Approach: They propose a Structurally Constrained Recurrent Network (SCRN) model and regularize it using existing techniques.
Outcome: The proposed model outperforms the LSTM model on non-English data while being much simpler.
Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are trained on extensive corpora to learn linguistic patterns, contextual nuances, and implicit elements of human values.
Approach: They propose to use word associations as low-level underlying representations to obtain a more robust picture of LLMs’ moral reasoning.
Outcome: The proposed method reveals detailed but systematic differences between LLMs and human associations.
MTLS: Making Texts into Linguistic Symbols (2024.emnlp-main)

Copied to clipboard

Challenge: In linguistics, all languages can be considered as symbolic systems . most work overlooks the properties of languages as symbol systems - aaron et al., 1989).
Approach: They propose a method to make texts into linguistic symbols to improve multilingual capability . they use a pre-training method to replace pre-trained language models with a vocabulary map .
Outcome: The proposed method improves multilingual capabilities on multilingual tasks using BERT and RoBERTa as the backbone.
Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: LVLMs are increasingly capable of responding in multiple languages . however, there is a lack of evaluation tools for LVLs that handle multiple languages.
Approach: They used an extended dataset in multiple languages to evaluate LVLMs' ability to generate explanations in multiple language combinations.
Outcome: The proposed dataset in multiple languages evaluates LVLMs' ability to generate explanations in other languages.
Features or Spurious Artifacts? Data-centric Baselines for Fair and Robust Hate Speech Detection (2022.naacl-main)

Copied to clipboard

Challenge: lexical biases in hate speech detection are limited when applied to real-world data, exhibiting limited out-of-distribution robustness and perpetuating harmful social biase.
Approach: They propose to disentangle spurious and authentic artifacts and analyze their impact on out-of-distribution fairness and robustness.
Outcome: The proposed models show that spurious artifacts require different treatments to attain robustness and fairness in hate speech detection.
PLATO-2: Towards Building an Open-Domain Chatbot via Curriculum Learning (2021.findings-acl)

Copied to clipboard

Challenge: PLATO-2 is a high-quality open-domain chatbot that can generate one-to-many mappings and improve response quality.
Approach: They propose a curriculum learning process to build a high-quality open-domain chatbot . they use a coarse-grained generation model and latent variables to train a generative model .
Outcome: The proposed model improves on Chinese and English data and can generate diverse responses and select the best response.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.
Translate & Fill: Improving Zero-Shot Multilingual Semantic Parsing with Synthetic Data (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve semantic parsing performance on target languages are limited.
Approach: They propose a Translate-and-Fill method that produces silver training data for a multilingual semantic parser.
Outcome: The proposed method produces silver training data for a multilingual parser.
A Neural Pairwise Ranking Model for Readability Assessment (2022.findings-acl)

Copied to clipboard

Challenge: Automatic Readability Assessment (ARA) is traditionally treated as a classification problem in NLP research.
Approach: They propose a neural ranking approach to automatic readability assessment (ARA) they propose 'neural' ranking methods that can be used to rank texts by reading level .
Outcome: The proposed approach performs well in monolingual single/cross corpus testing scenarios and achieves a zero-shot cross-lingual ranking accuracy of over 80% for both French and Spanish when trained on English data.
Multi-task and Multi-lingual Joint Learning of Neural Lexical Utterance Classification based on Partially-shared Modeling (C18-1)

Copied to clipboard

Challenge: Existing studies on multitask and multilingual joint learning focus on cross-task or cross-lingual knowledge transfer.
Approach: They propose to divide state-of-the-art neural lexical utterance classification into language-specific components that can be shared between different tasks and different languages.
Outcome: The proposed method is able to support multi-task and multi-lingual learning using Japanese and English data sets with three different lexical utterance classification tasks.
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages (2025.naacl-long)

Copied to clipboard

Challenge: In the brains of human bilinguals, syntax processing may occur in similar regions for their first and second language, depending on factors like when the second language was learned and language proficiency.
Approach: They propose to use sparse autoencoders to train Llama-3-8B and Aya-23-8B models to train multilingual models that share morphsyntactic representations of grammatical concepts.
Outcome: The proposed model can predict plural verbs in different languages by activating the same plural feature.
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: e.g., ChatGPT often provides inappropriate English-culture-related answers when users ask in non-English languages.
Approach: They build a benchmark of concrete and abstract cultural objects to evaluate the cultural dominance issue in large language models.
Outcome: The proposed model can significantly mitigate cultural dominance issue in large language models . the model can provide accurate answers in English, while the model is ethically sound .
DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational Recommendation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for conversational recommendation are limited to English and Chinese .
Approach: They propose a bilingual parallel human-to-human recommendation dialog dataset . the data item is annotated in two languages, both English and Chinese .
Outcome: The proposed dataset provides a testbed for future studies of multilingual and cross-lingual conversational recommendation.
Cross-lingual Transfer or Machine Translation? On Data Augmentation for Monolingual Semantic Textual Similarity (2024.lrec-main)

Copied to clipboard

Challenge: Using labeled NLI datasets for learning sentence embeddings leads to improved performance for natural language understanding tasks.
Approach: They compare two data augmentation techniques for learning better sentence embeddings . they use a cross-lingual transfer technique that exploits English resources as training data to yield non-English sentence embeds as zero-shot inference .
Outcome: The proposed techniques yield better performance on Japanese and Korean sentences.
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages (2022.emnlp-main)

Copied to clipboard

Challenge: Hate speech datasets focus on English-language content, hindering effective models . annotating hateful content is expensive, time-consuming and potentially harmful to annotators.
Approach: They propose to use ISO 639-1 codes to fine-tune models on one source language and apply them to another language.
Outcome: The proposed approach performs well on some tasks, but fails on many others.
Dissecting Span Identification Tasks with Performance Prediction (2020.emnlp-main)

Copied to clipboard

Challenge: Span identification tasks are a staple of applied NLP, but there is little insight on how their properties influence their difficulty.
Approach: They propose to build a model to predict span ID performance for unseen span ID tasks that can support architecture choices.
Outcome: The proposed model predicts span ID tasks for unseen span ID task in English, and the meta model predictable span ID performance.
Prompt-Tuning Can Be Much Better Than Fine-Tuning on Cross-lingual Understanding With Multilingual Language Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained multilingual language models show significant performance gains for zero-shot cross-lingual model transfer on a wide range of natural language understanding (NLU) tasks.
Approach: They do cross-lingual evaluation using prompt tuning and compare it with fine-tuning . prompt tuning achieves much better cross-linguistic transfer than fine- tuning .
Outcome: The results show that prompt tuning achieves better cross-lingual transfer than fine-tuning across datasets, with only 0.1% to 0.3% tuned parameters.
LangBridge: Multilingual Reasoning Without Multilingual Supervision (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to adapt language models for multilingual reasoning tasks require multilingual supervision.
Approach: They propose a zero-shot approach to adapt language models for multilingual reasoning tasks without multilingual supervision by bridging two models by introducing minimal trainable parameters between them.
Outcome: The proposed approach significantly improves multilingual reasoning capabilities on low-resource languages.
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages (2025.coling-main)

Copied to clipboard

Challenge: Question Answering datasets are scarce for languages other than English due to the cost and difficulties of collection and manual annotation.
Approach: They propose a method for generating and validating QA datasets for low-resource languages . they use English data as context to generate synthetic multiple-choice (MC) question-answer pairs .
Outcome: The proposed method maintains quality, reduces likelihood of factual errors, and circumvents costly annotation.
Construction of English-French Multimodal Affective Conversational Corpus from TV Dramas (L18-1)

Copied to clipboard

Challenge: Existing technologies for speech recognition and speech synthesis focus on non-verbal content and paralinguistic information.
Approach: They propose to construct a multimodal affective conversational corpus based on TV dramas . their data contain parallel English-French languages in lexical, acoustic, and facial features .
Outcome: The proposed corpus can be used to assess speech recognition, speech recognition and synthesis, linguistic, and paralinguistic speech-to-speech translation and multimodal dialog systems.
Localizing Open-Ontology QA Semantic Parsers in a Day Using Machine Translation (2020.emnlp-main)

Copied to clipboard

Challenge: a new toolkit for localizing a semantic parser for a language is proposed . the proposed approach is based on a method for question answering systems .
Approach: They propose a toolkit that leverages Neural Machine Translation systems to localize a semantic parser for a new language.
Outcome: The proposed approach outperforms state-of-the-art methods in 10 new languages . it can be deployed in restaurants and hotels in less than 24 hours .
Requirements and Motivations of Low-Resource Speech Synthesis for Language Revitalization (2022.acl-long)

Copied to clipboard

Challenge: Existing research on speech synthesis systems for three Indigenous languages in Canada requires tens of hours of audio recordings to be trained.
Approach: They build a system for three Indigenous languages spoken in Canada using 1 hour of training data and 10 hours of data to train low-resource models.
Outcome: The proposed system can produce speech with comparable naturalness to a Tacotron2 model trained with 10 hours of data.
Learn and Unlearn: Addressing Misinformation in Multilingual LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to unlearning large language models (LLMs) focus on English data, but they ignore multilingual contexts and can produce misleading, offensive, or otherwise fake content.
Approach: They investigate the propagation of information in multilingual large language models and evaluate unlearning methods to address harmful content in multi-lingual contexts.
Outcome: The proposed methods can effectively eliminate harmful content for all languages by addressing both English and the original language of the harmful data.
Demystifying Multilingual Reasoning in Process Reward Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study focuses on the use of large language models to solve multi-step reasoning tasks.
Approach: They propose to extend large language models to multilingual settings by extending process reward models to English . they train multilingual PRMs on a dataset spanning seven languages, which is translated from english .
Outcome: The proposed model improves accuracy and reduces early-stage reasoning errors.
SocioProbe: What, When, and Where Language Models Learn about Sociodemographics (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have outperformed other models on a wide range of tasks . however, there is still little understanding of their knowledge of higher-level aspects of language .
Approach: They investigate whether pre-trained language models have knowledge of sociodemographics . they use traditional probing techniques to probe the knowledge of single-GPU PLMs based on multiple English data sets .
Outcome: The results show that pre-trained language models outperform other models on a wide range of tasks.
Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods to improve Question Answering performance on non-English data are expensive and limited to evaluation set.
Approach: They propose a method to improve Question Answering performance without additional annotations by leveraging Question Generation models to produce synthetic samples in a cross-lingual fashion.
Outcome: The proposed method outperforms baselines on four datasets in English significantly . the proposed model outperformed baselines in english and is comparable to the validation set of the original SQuAD.
XL-WiC: A Multilingual Benchmark for Evaluating Semantic Contextualization (2020.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation benchmarks for assessing distinct meanings of words are tied to sense inventories, restricting their usage to knowledge-based representation techniques.
Approach: They propose a multilingual benchmark that models distinct meanings of words in English . they use a binary disambiguation task with gold standards in 12 new languages .
Outcome: The proposed model can model distinct meanings of words in English even when no tagged instances are available for a target language.
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) are mainly trained on English data and struggle with low-resource languages.
Approach: They propose to add a new language to Llama to improve classification accuracy for Persian tasks by aligning representations through bilingual pretraining and instruction datasets.
Outcome: The proposed model performs on generation and classification tasks with no adverse impact and sometimes even improvements on English tasks.
JAWAHER: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in instruction fine-tuning and alignment methods have enhanced the adaptability of large language models to user preferences.
Approach: They propose a benchmark to assess LLMs’ capacity to comprehend and interpret Arabic proverbs.
Outcome: The proposed model can generate accurate translations, but struggle to produce culturally nuanced and contextually relevant explanations.
English Contrastive Learning Can Learn Universal Cross-lingual Sentence Embeddings (2022.emnlp-main)

Copied to clipboard

Challenge: mSimCSE can learn high-quality universal cross-lingual sentence embeddings without any parallel data.
Approach: They propose a new language-based sentence embedding system that extends SimCSE to multilingual settings.
Outcome: The proposed method improves existing methods on retrieval and multilingual STS tasks.
To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging (2020.emnlp-main)

Copied to clipboard

Challenge: Using large amounts of unlabeled data to improve performance has become the foundation for many natural language processing tasks.
Approach: They propose a task-specific semi-supervised approach that uses unlabeled data in a more task-agnostic manner.
Outcome: The proposed approach achieves similar performance to BERT on a set of sequence tagging tasks with less financial and environmental impact.
CamemBERT: a Tasty French Language Model (2020.acl-main)

Copied to clipboard

Challenge: Pretrained language models are now ubiquitous in Natural Language Processing, but their use in other languages is limited.
Approach: They propose to train monolingual Transformer-based model for other languages using web crawled data instead of Wikipedia data and a relatively small web crawl dataset leads to better results.
Outcome: The proposed model performs as well as those obtained using larger datasets.
SAFARI: Cross-lingual Bias and Factuality Detection in News Media and News Articles (2024.findings-emnlp)

Copied to clipboard

Challenge: a new corpus of news media and articles is developed to assess political bias and factuality in cross-lingual contexts . integrity and objectivity of news are crucial in an age of information sharing across cultural and language landscapes - a recent study shows .
Approach: They propose a corpus of news media and articles for predicting political bias and factuality . they evaluate the cross-lingual ability of the models; however, they evaluate on English data .
Outcome: The proposed corpus is unprecedented in its collection and evaluates on English data.
Soft Gazetteers for Low-Resource Named Entity Recognition (2020.acl-main)

Copied to clipboard

Challenge: Existing named entity recognition models use gazetteers to improve performance, but they are limited in coverage and do not exist in low-resource languages.
Approach: They propose a method that integrates Wikipedia information into named entity models by cross-lingual entity linking.
Outcome: The proposed method improves on four low-resource languages with Wikipedia . it incorporates available information from english knowledge bases into neural models .
Cross-lingual Emotion Detection (2022.lrec-1)

Copied to clipboard

Challenge: Emotion detection is a useful tool for understanding human behavior, but constructing annotated datasets to train models can be expensive.
Approach: They propose to use English as the source language with Arabic and Spanish as target languages to train models for emotion detection in a target language.
Outcome: The proposed approaches surpass state-of-the-art models in Arabic and Spanish by 4% and 5% respectively.
Preference Tuning For Toxicity Mitigation Generalizes Across Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Detoxifying multilingual Large Language Models (LLMs) has become crucial due to their increasing global use.
Approach: They propose to use English preference tuning to study cross-lingual detoxification of LLMs.
Outcome: The proposed method reduces toxicity in multilingual LLMs by reducing the probability of mGPT-1.3B generating toxic continuations across 17 languages.
Language Agnostic Automatic Summarization Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: Existing evaluation methods for summarization of documents have been primarily focused on the English language.
Approach: They propose to use ROUGE and PYRAMID to evaluate non-English data using English and non- English data sets.
Outcome: The proposed evaluation methods can be adapted to non-English data, and the results show that they can perform well on non- English data.
Semantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual Transfer (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly incorporating multilingual capabilities, fueling the demand to transfer them into target language-specific models.
Approach: They propose a novel cross-lingual transfer technique that recycles embeddings from target language Pre-trained Language Models to transmit deep representational strengths to LLMs.
Outcome: The proposed technique outperforms existing methods in cross-lingual understanding setups and achieves faster convergence and lower loss during language adaptation.
Crosslingual Generalization through Multitask Finetuning (2023.acl-long)

Copied to clipboard

Challenge: Multitask prompted finetuning (MTF) has been shown to help large language models generalize to new tasks in a zero-shot setting, but so far explorations of MTF have focused on English data and models.
Approach: They apply multitask prompted finetuning to pretrained multilingual models and generate variants called BLOOMZ and mT0.
Outcome: The proposed models can generalize to non-English languages that have never been seen before.
Disentangling Language and Culture for Evaluating Multilingual Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Extensive evaluations of large language models (LLMs) are conducted on a wide range of models, revealing a notable cultural-linguistic synergy phenomenon, where models exhibit better performance when questions are culturally aligned with the language.
Approach: They propose a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of large language models by decomposing evaluation along dimensions of linguistic medium and cultural context.
Outcome: The proposed framework allows for a nuanced analysis of LLMs’ ability to process questions within both native and cross-cultural contexts cross-lingually.
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark (2026.findings-acl)

Copied to clipboard

Challenge: Authorship verification (AV) is a task of determining whether two texts were written by the same author.
Approach: They propose a benchmark for German AV comprising over 400k labeled text pairs.
Outcome: The proposed model outperforms baselines and state-of-the-art models by 0.09 and surpasses GPT-5 in a zero-shot setting by 0.08.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations